Marching Toward the Finals | Global Open-source AI Challenge Semi-Final Judging Officially Launches

On September 4, the semi-final judging for the three tracks of the Global Open-source AI Challenge --- "Agent Infra," "Boundless Agents," and "AI for Research" --- officially launches.

Previously, 120 projects stood out from 2,999 valid preliminary-round submissions. Among them, 30 projects from Agent Infra, 30 from Boundless Agents, and 40 from AI for Research advanced to the semi-finals; 20 projects from the Embodied Future track directly advanced to the finals according to the competition format. Following this round of semi-final judging, the Agent Infra and Boundless Agents tracks will each produce 15 advancing projects, and the AI for Research track will produce 20 advancing projects, which together with the 20 finalists from the Embodied Future track will form the top 70 of the finals.

The Three Tracks' Semi-Final Judging Has Different Priorities

Unlike the preliminary round, semi-final judges will focus on determining whether projects solve real problems, whether core pipelines can run completely, whether key capabilities can be independently verified, whether exceptions and risks can be properly handled, and whether results can be inspected, reproduced, and continuously iterated by third parties.

The three tracks have set differentiated judging criteria based on their respective technical directions and project forms:

Agent Infra Track: Verifying the Engineering Closed Loop of Multi-Agent Collaboration

This track focuses on next-generation agent infrastructure, examining multi-Agent collaboration, Skill engineering systems, and system engineering capabilities in complex enterprise task scenarios. Scoring consists of five parts: Scenario Value and Replicability (20%), Multi-Agent Collaboration Capability (25%), Skill Engineering System (20%), Engineering Implementation and Security Audit (30%), and Open-Source Contribution (5%).

Judging does not simply assess capability by Agent count or tool invocation frequency, but focuses on whether dynamic orchestration, shared state, human confirmation, exception recovery, access control, rollback, and auditing can jointly form a runnable and traceable enterprise-grade closed loop.

Boundless Agents Track: Verifying Task Value in Real Scenarios

This track targets industrial and lifestyle scenarios such as AI glasses, automotive, finance, education, and industrial manufacturing, focusing on answering: after agents enter an industry, what problems can they solve and what value can they create?

Semi-final judging revolves around Industry Scenario Value (25%), Agent Capability and Task Closed Loop (25%), Product Experience and Demo Completeness (20%), Technical Implementation Depth (15%), Security Compliance and Traceability (10%), and Open Reuse Contribution (5%).

Submissions need to demonstrate the complete pipeline from user requirement input, Agent understanding and decision-making, tool or knowledge base invocation, to result delivery, exception handling, and effectiveness verification. Meanwhile, data sources, usage authorization, privacy protection, risk boundaries, and engineering reproducibility are also important criteria. What judging seeks is not merely AI features that can be demonstrated, but product prototypes with clear user value, reliable running capability, and further replication potential.

AI for Research Track: Letting AI Enter the Scientific Discovery Process

This track focuses on the intersection of AI and scientific research, with Algorithm sub-track and Open Exploration sub-track, the two categories judged and ranked independently.

Algorithm sub-track is evaluated on four dimensions: Technical Performance (45%), Scientific Significance (30%), Method Innovation (20%), and Open-Source Contribution (5%); Open Exploration sub-track focuses on Problem Definition and Environment Design Quality (45%), Exploration Process and Scientific/Research Signals (35%), Inspectability and Sustainability (15%), and Open-Source Contribution (5%).

Judging not only focuses on metric performance, but also on whether scientific problems can be transformed into computable, verifiable, and reproducible research tasks. Anomaly discoveries with clear evidence, counterexamples, interpretable negative results, or effective revisions to problem definitions may likewise constitute valuable research signals. Reproducibility information such as code entry points, runtime environments, dependency versions, data sources, and expected outputs is an important basis for judging the credibility of results and potential for follow-up research.

Semi-Final Online Defense: Independent Judging, Fair and Professional

Online defense follows a unified process: each team completes project presentation and judge Q&A within the allotted time, and judges score independently. The defense centers on evidence clarification based on submitted materials, and newly added features on-site do not substitute for formal results submitted before the deadline. For specific arrangements, please visit the GOAI official website at goaihz.com.

Judges independently pre-review, ask questions, and score; comments must correspond to specific evidence and include improvement suggestions. For projects with large score discrepancies among judges, near the advancement line, or with questionable facts, the organizing committee will initiate a review procedure; network interruptions, lateness, absence, and other situations are recorded and handled according to unified rules. Matters involving missing materials, data authorization, result credibility, or conflicts of interest will be verified separately. Judges must strictly maintain the confidentiality of submission materials, judging discussions, and results, and should proactively request recusal when a conflict of interest exists.

According to venue configuration, each online venue is in principle composed of approximately 3 expert judges, 1 host and timekeeper, and 1 judging recorder, with a central coordination lead overseeing the progress and emergencies of all venues.

Semi-Final Judging Schedule

· Sep. 4: Agent Infra Track Online Defense

· Sep. 5--6: Boundless Agents and AI for Research Tracks Hold Simultaneous Defenses in Two Online Venues

· Sep. 7--9: Complete Judging Review and Calibration

· Sep. 10: Finalists Announcement

The Top 70 Finalists Are About to Be Determined, Marching Toward the Hangzhou Finals

From nearly 3,000 preliminary-round submissions, to 120 advancing projects, to 50 final spots across the three tracks, the competition's evaluation is deepening along a clear path: from whether a problem is worth solving and whether a technical solution is viable, to whether a system can run, whether results can be verified, and whether outcomes can be openly reused.

After the semi-finals, projects that truly advance to the finals need to prove not only that their ideas are novel enough, but also that their technology can be implemented, their value can withstand scrutiny, and their outcomes have the potential for continuous evolution.

On September 10, the top 50 finalists from the three tracks will be officially announced. At that time, the advancing teams will join the top 20 from the Embodied Future track, gathering in Hangzhou from September 22 to 23 to participate in the Finals & GOAI DAY of Global Open-source AI Challenge (GOAI).

The semi-final judging of the Global Open-source AI Challenge officially launches. Who will earn the ticket to the Hangzhou finals? Stay tuned.